Papers with continuous reward modeling
ARF-RLHF: Adaptive Reward-Following for RLHF through Emotion-Driven Self-Supervision and Trace-Biased Dynamic Optimization (2026.acl-long)
Copied to clipboard
| Challenge: | prevailing RLHF methods such as PPO and DPO depend on large-scale binary preference annotations. |
| Approach: | They propose a method which converts natural feedback into continuous preference trajectories and optimizes them using the novel TraceBias algorithm. |
| Outcome: | The proposed approach outperforms PPO and DPO in a variety of domains and improves alignment by up to 7.6% across diverse LLMs and preference domains. |